Papers by Pedro Ortiz Suarez
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing web crawling pipelines are used to collect large corpora raw data, but the main way to collect such data is through manual data extraction. |
| Approach: | They propose to use a web crawler to extract and classify data from a multilingual web corpus and an automated annotation pipeline to improve it. |
| Outcome: | The proposed version of OSCAR could be used to pre-train large generative language models and other applications in Natural Language Processing and Digital Humanities. |
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)
Copied to clipboard
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot
| Challenge: | Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available. |
| Approach: | They propose to use a contextualised language model to analyse historical states of language in French. |
| Outcome: | The proposed model is based on a corpus of historical texts and is evaluated with an NLP task. |
A Data-driven Approach to Named Entity Recognition for Early Modern French (2022.coling-1)
Copied to clipboard
| Challenge: | Named entity recognition is an important task in natural language processing. |
| Approach: | They propose to use a data-driven approach to identify historical French with fine-grained annotations instead of a specialised architecture to tackle particularities. |
| Outcome: | The proposed corpus is larger than the most popular NER evaluation corpora for both Contemporary English and French. |
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus (2025.findings-acl)
Copied to clipboard
Matthieu Futeral, Armel Randy Zebaze, Pedro Ortiz Suarez, Julien Abadji, Rémi Lacroix, Cordelia Schmid, Rachel Bawden, Benoît Sagot
| Challenge: | Existing studies show that multimodal large language models can learn from text-image data. |
| Approach: | They propose to train multimodal large language models on large amounts of text-image data . they also show a boost in few-shot learning performance across various multilingual tasks . |
| Outcome: | The proposed dataset is not public and is only in English . it is the first large-scale multilingual and multimodal document corpus crawled from the web. |
BERTrade: Using Contextual Embeddings to Parse Old French (2022.lrec-1)
Copied to clipboard
| Challenge: | a growing interest in digital humanities for automatic processing and annotation of historical texts is generating new models for historical languages. |
| Approach: | They use POS-tagging and dependency parsing to evaluate contextual word embedding models . Old French is one of the historical languages for which they have the largest amount of syntactically annotated data . |
| Outcome: | The proposed model can be used to improve performance in Old French, the authors show . they use POS-tagging and dependency parsing to evaluate the model's quality . |
Tokenizer Choice For LLM Training: Negligible or Crucial? (2024.findings-naacl)
Copied to clipboard
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Buschhoff, Charvi Jain, Alexander Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, Nicolas Flores-Herr
| Challenge: | Recent success of large language models has been driven by curating the training dataset composition, scaling of model architectures and advancements in pretraining objectives, leaving tokenizer influence as a blind spot. |
| Approach: | They conduct a comprehensive study on the influence of tokenizer choice on LLM downstream performance by training 24 mono- and multilingual LLMs at a 2.6B parameter scale. |
| Outcome: | The proposed model can significantly impact the model's downstream performance and training costs. |
Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a new pipeline can be used to create corpora for over-looked languages . |
| Approach: | We propose a new pipeline that can filter a single snapshot in twohours. |
| Outcome: | The proposed pipeline can filter a single snapshot in twohours. |
A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages (2024.lrec-main)
Copied to clipboard
Jorge Palomar-Giner, Jose Javier Saiz, Ferran Espuña, Mario Mina, Severino Da Dalt, Joan Llop, Malte Ostendorff, Pedro Ortiz Suarez, Georg Rehm, Aitor Gonzalez-Agirre, Marta Villegas
| Challenge: | CATalog 1.0 is the largest text corpus in Catalan to date . CURATE is a pipeline that can be parallelizable to run in high performance clusters . |
| Approach: | They propose a data pipeline that uses binary filters to filter documents based on text quality . they optimised the pipeline to run in high performance clusters . |
| Outcome: | The proposed pipeline is optimized for high performance cluster environments and runs in high performance. |